Papers with statistical method
Mandarinograd: A Chinese Collection of Winograd Schemas (2020.lrec-1)
Copied to clipboard
| Challenge: | Mandarinograd is a corpus of Winograd Schemas in Mandarin Chinese . WS are hard to collect and few datasets are publicly available . |
| Approach: | They introduce a corpus of Winograd Schemas in Mandarin Chinese . they describe the difficulties faced when building the corpus and explain how they overcome the anomalies. |
| Outcome: | The proposed corpus of Winograd Schemas in Mandarin Chinese is hard to build and resistant to statistical methods. |
uniblock: Scoring and Filtering Corpus with Unicode Block Information (D19-1)
Copied to clipboard
| Challenge: | Existing methods to remove sentences consisting of illegal characters are tedious and repetitive. |
| Approach: | They propose a statistical method to identify illegal characters in natural language processing . they use a fixed-size feature vector to generate a Gaussian mixture model for each sentence . |
| Outcome: | The proposed method can score sentences and filter corpus on clean corpus and improve performance. |
Stubborn Lexical Bias in Data and Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent work has focused on spurious correlations between features and labels in training data . but, we find strong evidence of corresponding bias in the trained models . |
| Approach: | They propose a method to reduce spurious correlations in training data by reweighting it using a large pool of extracted features. |
| Outcome: | The proposed method reduces spurious correlations in training data, but still finds strong evidence of bias in trained models. |
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants (2024.lrec-main)
Copied to clipboard
Chloe Sekkat, Fanny Leroy, Salima Mdhaffar, Blake Perry Smith, Yannick Estève, Joseph Dureau, Alice Coucke
| Challenge: | Recent studies show voice assistants do not perform equally well for everyone . however, research on demographic robustness of speech technologies is still scarce . |
| Approach: | They propose a statistical method to detect demographic bias using a large dataset with controlled demographic tags. |
| Outcome: | The proposed method shows statistically significant differences in performance across age, dialectal region and ethnicity. |